feat(bench): v3 conditions, Wright Agent Score, and agent adapters - #464
Merged
Merged
Conversation
…le canary Refs #414. Adds adapters for pi and the Devin CLI that report per-turn usage, transcript, and loaded context, with host configuration kept out of the run (no context files or extensions for pi; an isolated HOME, no cross-tool rules, and denied MCP tools for Devin). Runs now fail a canary when instruction files exist above the workspace, the default output directory moves outside the repository, and results record whether network off was checked. Updates the matrix example and the benchmark contract.
Refs #414. wiki-snapshot fetches the Markdown mirror once into a local directory with per-document hashes and a snapshot hash. Runs with knowledge wiki require a snapshot and record its identity. Fetching goes through curl because the mirror returns 403 to Python's HTTP client.
…ll and level Refs #414. Local work, not pushed. wiki-snapshot now crawls category pages (the manifest lists only the first upstream page), with retries, resume, and bounded concurrency. wiki-skill builds a progressive-disclosure workshop-wiki skill from a snapshot, with OverPy spellings taken from the workshop-rs catalog and the opy-rs manifest and kept only when the pinned upstream compiler contains them. Adds the wiki-skill knowledge level, adapter support for a second skill, and line-buffered usage files so a killed run keeps its usage.
Reject changed snapshot and skill content before evaluation, correct the generated license notice, and align documentation and pilot matrices with the wiki-skill condition. Refs #414
Run Codex and Antigravity with isolated configuration and normalized usage, isolate pi authentication, and restrict trial writes with the macOS sandbox. Keep unobserved context explicit and exclude provider failures from agent outcome metrics. Refs #414
Clear stale pi provider errors after a successful response and classify observed transport failures as infrastructure exits. Refs #414.
Restrict outside-workspace instruction reads in the macOS file sandbox. Verify workspace instructions remain readable and host rules are denied before rerunning Devin pilot gaps. Refs #414.
Codex's workspace-write seatbelt cannot be applied inside the harness sandbox-exec profile (sandbox_apply: Operation not permitted), so every Codex tool write was denied in the preflight. Run Codex with its own sandbox off and let the harness file sandbox bound writes. Network off is declared-only for this adapter.
…guage Refs #467. Adds greenfield-opy-elimination-race, understand-opy-events, modify-opy-kill-hud, modify-opy-extract-subroutine, and repair-opy-self-kill-score, plus Workshop twins of the last four (compiled from the OverPy references by the pinned compiler). All are test split, so each language track now has 8 held-out scenarios. Every scenario has a reference that passes, a seed that fails, and negatives that fail exactly the declared checks. Updates the ana-paintball hallucinated-name negative, which Wright now rejects like the upstream compiler.
…ght Agent Score Refs #466, #467. Conditions become tool (none, wright, overpy) x skills (wright, workshop, opy, workshop-format) x knowledge x network, labelled tool[+skill...]/knowledge/network. The overpy tool and the language skills apply only to scenarios of their language, and matrix skips the rest. The harness adds no text to the scenario prompt. Results record protocol, Wright binary hash, skill hashes, suite hash, harness commit, agent info, and a status. usable now also blocks on unsafe edits and unavailable required graders. The wiki is copied into the workspace instead of linked. matrix stops after repeated provider interruptions and returns 3. New score command computes per-language scenario macro-average scores with a clustered bootstrap interval, Pass^k, and a score card. report pairs against any reference condition, splits by language, and uses clustered intervals.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #414, #466, #467. Brings the agent benchmark to the v3 contract: language-appropriate controls, a score card, and the adapters and wiki inputs it needs.
Conditions (v3). A cell is
tool[+skills]/knowledge/network.toolisnone,wright, oroverpy(OverPy scenarios only); skills arewright-skill,workshop-skill,opy-skill,workshop-format-skill; knowledge isnone,wiki(copied into the workspace), orweb. Pairs not applicable to a scenario's language are skipped. The harness no longer adds text to the task prompt.Contract
wright-agent-bench/v3.status(completed,provider-interrupted,timeout,agent-error,invalid),usablewithusableReason(lint errors, unsafe edits, and an unavailable grader block it),protocol,agentInfo, suite identity (version and hash), skill identities, harness commit, per-tooltoolUse, andnetworkEnforcementdisclosure. Canaries fail a run when a tool outside the condition is reachable.Score
wright-agent-score/v1.agent_bench.py scoregives one card per language track on the canonical cellwright+wright-skill/none/off: scenario macro-average, two-stage bootstrap interval, Pass^k, exclusions, provisional reasons, and a refusal when runs come from different environments.reportgains clustered intervals and--reference.Scenarios. Held-out OPY and Workshop scenarios toward 8 per language, plus a negative for hallucinated names.
Adapters and wiki. pi, Devin, native Codex, and Antigravity adapters; a pinned wiki snapshot crawled by category and a local
workshop-wikiskill build. Wiki content stays local and is not committed.Verification. 64 unit tests pass with the pinned oracle installed (
python3 -m unittest discover -s benchmarks/agent).Not verified. Network
offis declared, not enforced, and the card says so. No full-matrix run has been done on v3; the scenario count per language has not been checked against the 8 target.